Back

Brain Informatics

Springer Science and Business Media LLC

All preprints, ranked by how well they match Brain Informatics's content profile, based on 10 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
Machine Learning for Longitudinal Brain-Age Prediction

Wegmann, M.; Ganz, M.; Svensson, J. E.; Plaven-Sigray, P.; Dörfel, R. P.

2025-12-11 health informatics 10.64898/2025.12.10.25341964 medRxiv
Top 0.1%
11.9%
Show abstract

Cross-sectional brain age models have demonstrated high accuracy and reliability for predicting chronological age based on structural brain features derived from single MRI scans. However, these models cannot separate baseline variation from true aging-related changes or noise. Longitudinal models address this limitation by predicting inter-scan intervals from paired MRI scans, controlling for baseline factors through repeated measurements. Using OASIS-3 data, we compare a cross-sectional 3D CNN against three longitudinal architectures for predicting inter-scan intervals: LILAC (Siamese neural network), LILAC+ (enhanced Siamese network with multi-layer perceptron), and AM (variational autoencoder). Longitudinal models substantially outperformed the cross-sectional approach, with LILAC+ achieving best performance (MSE = 1.97 years2, MAE = 0.99 years, r = 0.86, R2 = 0.71). Our results suggest that direct modeling of longitudinal change is more effective at capturing individual aging trajectories than deriving intervals from cross-sectional predictions.

2
Investigating the Data Addition Dilemma in Longitudinal TBI MRI

Titikhsha, A.; Akhtar, M.; Mollah, A. M.

2025-09-30 health informatics 10.1101/2025.09.29.25336939 medRxiv
Top 0.1%
8.1%
Show abstract

Clinical machine learning (CML)for brain MRI often assumes that more data guarantees better performance, yet added samples can reduce accuracy when they arise from a different distribution, a phenomenon known as the Data Addition Dilemma. We present a systematic study of this issue in longitudinal TBI MRI, where acute baseline scans (S1) and follow-up scans (S2) differ substantially. Using a 14-subject, 28-scan cohort, we quantify the combined effects of intra-subject session shifts and inter-subject variability on severity classification. We evaluate four training schemes: (1) intra-session upper bound (S1[->]S1), (2) cross-session OOD testing (S1[->]S2), (3) pooled training (S1+S2[->]S1,S2), and (4) LOSO-IPA, which adds one unlabeled S2 scan per patient. With a lightweight logistic-regression model on PCA features, we show that naive pooling can degrade accuracy, pooled training trades baseline performance for modest robustness gains, and LOSOIPA recovers accuracy close to the intra-session limit. We recommend per-subject follow-up anchoring and diagonal CORAL alignment to mitigate session effects. These results clarify when additional data help or hinder CML workflows and provide a minimally invasive strategy for reliable longitudinal TBI severity assessment.

3
Geometric Deep Learning Methods for Improved Generalizability in Medical Computer Vision: Hyperbolic Convolutional Neural Networks in Multi-Modality Neuroimaging

Ayubcha, C.; Sajed, S.; Omara, C.; Singh, S. B.; Lokesha, Y. U.; Liu, A.; Aziz-Sultan, M. A.; Smith, T. R.; Beam, A.

2024-10-14 health informatics 10.1101/2024.10.12.24315391 medRxiv
Top 0.1%
7.2%
Show abstract

ObjectiveThis study investigates the potential advantages of hyperbolic convolutional neural networks (HCNNs) over traditional convolutional neural networks (CNNs) in neuroimaging tasks. Materials and MethodsWe conducted a comparative analysis of HCNNs and CNNs across various medical imaging modalities and diseases, with a focus on a compiled multi-modality neuroimaging dataset. The models were assessed for performance parity, robustness to adversarial attacks, semantic organization of embedding spaces, and generalizability. Zero-shot evaluations were also performed with ischemic stroke non-contrast CT images. ResultsHCNNs matched CNN performance on less complex settings and demonstrated superior semantic organization, and robustness to adversarial attacks. While HCNNs equaled CNNs in out-of-sample datasets identifying Alzheimers disease, in zero-shot evaluations, HCNNs outperformed CNNs and radiologists. DiscussionHCNNs deliver enhanced robustness and organization in the neuroimaging data. This likely underlies why while HCNNs perform similarly to CNNs with respect to in-sample tasks, they confer improved generalizability. Nevertheless, HCNNs encounter efficiency and performance challenges with larger, complex datasets. These limitations underline the need for further optimization of HCNN architectures. ConclusionHCNNs present promising improvements in generalizability and resilience for medical imaging applications, particularly in neuroimaging. Despite challenges with larger datasets, HCNNs enhance performance under adversarial conditions and offer better semantic organization, suggesting valuable potential in generalizable deep learning models in medical imaging and neuroimaging diagnostics.

4
WHITE-Net : White matter HyperIntensities Tissue Extraction using deep learning Network

Cathala, C.; Kherif, F.; Thiran, J.-P.; Bussy, A.; Draganski, B.

2025-01-09 health informatics 10.1101/2025.01.09.25320242 medRxiv
Top 0.1%
6.8%
Show abstract

Given the high prevalence of aging-associated cerebral small vessel disease in the general population, accurate detection of the related white matter hyperintensities (WMH) in large-scale magnetic resonance imaging (MRI) studies is of critical importance. The performance of currently available semi-automated and automated methods for WMH classification is hampered by their inherent dependence on MRI contrast parameters and long computational processing time. We sought to improve the accuracy and computational cost of automated WMH detection by creating a whole-brain deep learning-based framework: WHITE-Net. We use a 3D ResUNet architecture trained on manually segmented WMHs from fluid-attenuated inversion recovery MRI (n=141) and test its accuracy in a large-scale dataset (n=192). We demonstrate a good generalizability across WMH lesion loads, different MRI scanner vendors, field strengths, imaging protocols, and MR contrasts. The comparison to existing WMH segmentation tools shows a similar to superior accuracy performance at significantly lower computational cost. WHITE-Net tool performance makes it well-suited for application to large-scale MRI datasets, enabling the study of the aging brain while offering the advantage of detecting early or subtle WMH changes often missed by other methods.

5
Enabling Advanced Multi-Modal Neuroimaging Analysis within a Trusted Research Environment

Hotchkiss, L.; Squires, E.; Gallacher, J.; Morris, C.; Newbury, M.; Lyons, R.; Thompson, S.

2024-02-14 health informatics 10.1101/2024.02.13.24302751 medRxiv
Top 0.1%
6.6%
Show abstract

IntroductionGlobally, 55 million individuals have dementia, with an increasing annual incident of 10 million. Enabling development of new multi-modal models can improve the current diagnostic pathways and potentially contribute to early diagnosis and treatment of dementia. Here, we report how multi-modal resources is achieved within the successful Trusted Research Environment (TRE) providing access to 60+ cohort datasets for dementia research, the Dementias Platform UK (DPUK). ObjectivesWe aimed to identify the challenges of the storage, distribution and analysis of neuroimaging data and how we could implement a comprehensive infrastructure to deal with these. The problems we specifically aimed to address were how to: anonymise scans, store large amounts of data, standardise datasets to a common format, extract metadata, provision the data, and allow for analysis. MethodsWhile, data within majority of existing research platforms are focused on a single aspect, DPUK data provides an enriched view of disease dynamic for dementia cohorts by providing access to linkable brain imaging and genomic data at the individual-level. We document various stages and capacities required for multi-modal neuroimaging analysis for dementia and conclude that achieving research ready assets to enable neuroimaging analysis for dementia from existing resources requires an engineered process to facilitate multiple aspects of curation, provisioning and large scale analysis. ResultsWe developed an ingest pipeline for neuroimaging data to meet the requirements set out in the objectives. This involved standardising all datasets to the Brain Imaging Data Structure, defacing scans and anonymising data, using MinIO for data storage and extracting metadata from header information for data discovery and provisioning. ConclusionThe neuroimaging ingest pipeline developed has allowed for the distribution of imaging datasets within DPUK which has facilitated multi-modal research on anonymised and standardised data. Our pipelines create research-ready datasets in a simplified way, reducing the time and effort of getting these datasets ready for data sharing and making the process easier for the data owners.

6
Selectively Augmented Decision Tree for Explainable Dementia Detection

Kamalov, F.; Thabtah, F.; Peebles, D.; Ibrahim, A.

2026-02-04 health informatics 10.64898/2026.02.03.26345441 medRxiv
Top 0.1%
6.0%
Show abstract

Timely and accurate diagnosis of dementia remains a critical yet challenging task. Although machine learning (ML) techniques have shown considerable promise in dementia detection, their inherent complexity often results in opaque, "black-box" models that limit clinical acceptance and usability. In this paper, we propose a Selectively Augmented Decision Tree (SADT), an interpretable AI model specifically designed for dementia detection. SADT incorporates a structured three-phase pipeline consisting of feature selection, data balancing, and construction of a transparent decision tree classifier. We apply SADT to the OASIS dataset and evaluate it empirically, showing that SADT outperforms traditional ML benchmarks, validating its effectiveness. In addition to its superior performance, SADT also mirrors aspects of human decision-making in its sequential, rule-based prioritization of key features. This approach aligns with cognitive models of cue use and heuristic reasoning, making it not only clinically transparent but also psychologically aligned with how diagnostic decisions are often made in practice. SADTs strong predictive performance and interpretability grounded in human reasoning facilitates explanation and human scrutiny, and has the potential to improve both clinical decision-making and trust in AI-assisted diagnosis.

7
Auxiliary Clinical Prompt Integration into Vision-Language Prompt SAM for Brain Tumor Segmentation

Hakata, Y.; Oikawa, M.; Fujisawa, S.

2026-04-17 health informatics 10.64898/2026.04.15.26351001 medRxiv
Top 0.1%
5.5%
Show abstract

BackgroundAdult diffuse glioma is a representative class of primary brain tumors for which accurate MRI-based tumor segmentation is indispensable for treatment planning. Conventional automated segmentation methods have relied primarily on image information and spatial prompts, and auxiliary clinical information that is routinely acquired in clinical practice has not been sufficiently exploited as an input. ObjectiveBuilding on a dual-prompt-driven Segment Anything Model (SAM) extension framework [20] that fuses visual and language reference prompts, we propose a method that integrates patient demographics, unsupervised molecular cluster variables derived from TCGA high-throughput profiling, and histopathological parameters as learnable prompt embeddings, and we evaluate its effect on the accuracy of lower-grade glioma (LGG) MRI segmentation. MethodsAn auxiliary prompt encoder converts clinical metadata into high-dimensional embeddings that are fused with the prompt representations of Segment Anything Model (SAM) ViT-B through a cross-attention fusion mechanism. The TCGA-LGG MRI Segmentation dataset (Kaggle release by Buda et al. [24]; n = 110 patients; WHO grade II-III) was split at the patient level (train/val/test = 71/17/22) using three different random seeds, and the three slices with the largest tumor area were extracted from each patient. To avoid pseudo-replication arising from multiple slices per patient and repeated measurements across seeds, our primary analysis aggregated Dice and 95th-percentile Hausdorff distance (HD95) to the patient x seed unit (n = 66); secondary analyses at the unique-patient level (n = 22) and at the per-slice level (n = 198) are also reported. Pairwise comparisons used paired t-tests with Bonferroni correction (k = 3) and Wilcoxon signed-rank tests, and a permutation test (K = 30) served as an auxiliary check of effective use of the auxiliary information. ResultsAt the patient x seed level (n = 66), Proposed (full clinical) achieved a Dice gain of {Delta} = +0.287 over the zero-shot SAM ViT-B baseline (paired-t p = 4.2 x 10-{superscript 1}, Cohens d_z = +1.25, Bonferroni-corrected p << 0.001; Wilcoxon p = 2.0 x 10-{superscript 1}), and HD95 improved from 218.2 to 64.6. Because zero-shot SAM is not designed for domain-specific medical segmentation, the large absolute HD95 gap largely reflects the expected domain gap rather than a competitive baseline. The additional contribution of the full clinical configuration over the demographics-only configuration was {Delta} Dice = +0.023 (paired-t p = 0.057, Bonferroni-corrected p = 0.172), which did not reach statistical significance at the patient level and is reported as a directional trend. The permutation test (K = 30, seed 2025) yielded real-metadata Dice = 0.819 versus a shuffled-metadata mean of 0.773, giving an empirical p = 0.032 = 1/(K + 1), which is at the resolution limit of this test and should therefore be interpreted as preliminary evidence. ConclusionsIntegrating auxiliary clinical information as multimodal prompts produced a large improvement over the zero-shot SAM baseline on this LGG cohort. More importantly, a robustness analysis showed that Proposed (full clinical) outperformed the trained Base (no auxiliary information) under all tested spatial-prompt conditions, including perfect centroid ({Delta} = +0.014), and that the advantage was most pronounced in the prompt-free regime ({Delta} = +0.231, p = 0.039), where the base model collapsed but the proposed model maintained meaningful segmentation by leveraging clinical metadata alone. The additional contribution of molecular and histopathological information beyond demographics was not statistically resolved at the patient level ({Delta} = +0.023, n.s.). Establishing clinical utility will require external validation on larger multi-center cohorts and direct comparisons with established segmentation methods.

8
BrainGT: Multifunctional Brain Graph Transformer for Brain Disorder Diagnosis

Shehzad, A.; Zhang, D.; Xia, F.; Yu, S.; Abid, S.; Cheng, X.; Zhou, J.

2024-08-31 health informatics 10.1101/2024.08.30.24312819 medRxiv
Top 0.1%
5.5%
Show abstract

Functional brain networks play an essential role in the diagnosis of brain disorders by enabling the identification of abnormal patterns and connections in brain activities. Previous methods often rely on whole brain functional connectivity approaches to construct these networks using Functional Magnetic Resonance Imaging (fMRI) data. However, these approaches introduce noise and overlook localized disruptions within specific brain subnetworks, leading to potential misdiagnoses. To address this challenging issue, we propose mBrainGT, a modular brain graph transformer model that focuses on modular functional connectivity (mFC) to improve the diagnosis of brain disorders. Compared to existing methods, mBrainGT constructs and analyses functional brain subnetworks individually, reflecting the inherent structure of the brain. It captures both local features within each modular network and their interactions through self-attention and cross-attention mechanisms. It also learns global interactions via adaptive fusion. We validate mBrainGT on three benchmark datasets (ADNI, PPMI, and ABIDE). The results demonstrate that mBrainGT outperforms existing methods in diagnostic accuracy, providing more robust and precise representations of the brain network essential for accurate disease detection. Our study highlights the potential of modular connectivity-based graph learning in the refinement of brain disorder diagnostics, offering a more precise and biologically relevant representation of functional brain networks.

9
Continuous lesion images drive more accurate predictions of outcomes after stroke than binary lesion images

Hope, T. M. H.; Neville, D.; Seghier, M. L.; Price, C. J.

2024-10-07 neuroscience 10.1101/2024.10.04.616726 medRxiv
Top 0.1%
5.5%
Show abstract

Current medicine cannot confidently predict who will recover from post-stroke impairments. Researchers have sought to bridge this gap by treating the post-stroke prognostic problem as a machine learning problem. Consistent with the observation that these impairments are caused by the brain damage that stroke survivors suffer, information concerning where and how much lesion damage they have suffered conveys useful prognostic information for these models. Much recent research has considered how best to encode this lesion information, to maximise its prognostic value. Here, we consider an encoding that, while not novel, has never before been formally examined in this context: continuous lesion images, which encode continuous evidence for the presence of a lesion, both within and beyond what might otherwise be considered the boundary of a binary lesion image. Current state of the art models employ information derived from binary lesion images. Here, we show that those models are significantly improved (i.e., with smaller Mean Squared Error between predicted and empirical outcome scores) when using continuous lesion images to predict a wide range of cognitive and language outcomes from a very large sample of stroke patients. We use further model comparisons to locate the predictive advantage to the provision of continuous lesion evidence beyond the boundary of binary lesion images. The continuous lesion images thus provide a straightforward way to incorporate details of both lesioned and non-lesioned tissue when predicting outcomes after stroke.

10
Comparison of Explainable AI Models for MRI-based Alzheimer's Disease Classification

Chattopadhyay, T.; Joshy, N. A.; Jagad, C.; Gleave, E.; Thomopoulos, S. I.; Feng, Y.; Villalon-Reina, J. E.; Laltoo, E.; Joshi, H.; Venkatasubramanian, G.; John, J. P.; Steeg, G. V.; Ambite, J. L.; Thompson, P. M.

2024-09-17 neuroscience 10.1101/2024.09.17.613560 medRxiv
Top 0.1%
5.4%
Show abstract

Deep learning models based on convolutional neural networks (CNNs) have been used to classify Alzheimers disease or infer dementia severity from 3D T1-weighted brain MRI scans. Here, we examine the value of adding occlusion sensitivity analysis (OSA) and gradient-weighted class activation mapping (Grad-CAM) to these models to make the results more interpretable. Much research in this area focuses on specific datasets such as the Alzheimers Disease Neuroimaging Initiative (ADNI) or National Alzheimers Coordinating Center (NACC), which assess people of North American, predominantly European ancestry, so we examine how well models trained on these data generalize to a new population dataset from India (NIMHANS cohort). We also evaluate the benefit of using a combined dataset to train the CNN models. Our experiments show feature localization consistent with knowledge of AD from other methods. OSA and Grad-CAM resolve features at different scales to help interpret diagnostic inferences made by CNNs.

11
Predicting cognitive impairment using novel functional features of spatial proximity and circularity in the digital clock drawing test

Pinheiro, A.; Karjadi, C.; Tripodis, Y.; Kolachalama, V. B.; Lunetta, K. L.; Demissie, S.; Liu, C.; Au, R.; Mohammed, S.

2026-03-16 health informatics 10.64898/2026.03.14.26348336 medRxiv
Top 0.1%
5.4%
Show abstract

The digital clock drawing test (dCDT) is a cognitive screening tool employing a digital pen. While many studies rely on summary statistics of dCDT features to predict cognitive outcomes, these approaches often involve subjective decisions such as feature selection and imputation. In this study, we introduce novel dCDT features, expressed as mathematical functions, and compare them to commonly used summary features. We included dCDTs from 3,415 participants from the Framingham Heart Study. Random forest models with five-fold cross-validation were trained to distinguish participants with mild cognitive impairment or dementia from cognitively intact participants. When combined with established time-based features, functional features related to spatial proximity and circularity demonstrated predictive power comparable to commonly used summary features. Our findings highlight the potential of integrating functional features to detect subtle motions and behaviors in digital cognitive assessments, offering new tools that may enhance diagnostic accuracy and support early detection strategies.

12
AI-based model for T1-weighted brain MRI diagnoses Amyotrophic Lateral Sclerosis

Turrisi, R.; Forzanini, F.; Stanziano, M.; Nigri, A.; Fedeli, D.; Giovanna, C.; Laura, L.; Manera, U.; Moglia, C.; Valentini, M. C.; Calvo, A.; Chio', A.; Barla, A.

2024-04-28 health informatics 10.1101/2024.04.26.24306438 medRxiv
Top 0.1%
5.4%
Show abstract

Amyotrophic Lateral Sclerosis (ALS) is an incurable deadly motor neuron disease that causes the gradual deterioration of nerve cells in the spinal cord and brain. It impacts voluntary limb control and can result in breathing impairment. ALS diagnosis is often challenging due to its symptoms overlapping with other medical conditions and many tests must be performed to rule out other conditions, as easily identifiable biomarkers are still lacking. In this study, we explore T1-weighted (T1w) brain Magnetic Resonance Imaging (MRI), a non-invasive neuroimaging approach which has shown to be a reliable biomarker in many medical fields. Nonetheless, current literature on ALS diagnosis fails to retrieve evidence on how to identify biomarkers from T1w MRI. In this paper, we leverage Artificial Intelligence (AI) methods to unveil the unexplored potential of T1w brain MRI for distinguishing ALS patients from those who have similar symptoms but different diseases (mimicking). We consider a retrospective single-center dataset of brain T1-weighted MRIs collected from 2010 to 2018 recruited from the Piemonte and Valle dAosta ALS register (PARALS). The collection includes 548 patients with ALS and 106 with mimicking diseases. Our goal is to develop and validate a ML diagnostic model based exclusively on T1w MRI distinguishing the two classes. First, we extract a set of radiomic features and two sets of Deep Learning (DL)-based features from MRI scans. Then, using each representation, we train 8 binary classifiers. The best results were obtained by combining DL-based features with SVM classifier, reaching an F1-score of 0.91, and a Precision of 0.88, a Recall of 0.94, and an AUC of 0.7 considering the ALS group as the positive class in the testing set.

13
Deep generative models for vessel segmentation in CT angiography of the brain

van Voorst, H.; Su, J.; Konduri, P.; Majoie, C.; Roos, Y.; Emmer, B.; Marquering, H.; de Vos, B.; Caan, M.; Isgum, I.

2025-03-19 health informatics 10.1101/2025.03.07.25322919 medRxiv
Top 0.1%
5.4%
Show abstract

Automated vessel segmentation in brain CT angiography (CTA) remains challenging despite the potential benefit of applications. Expert acquisition of reference vessel segmentations is a laborious task. We propose an unsupervised generative deep learning approach that can be trained for vessel segmentation in brain CTA using a large dataset (n=908) of unlabelled brain CTAs and non-contrast enhanced CTs (NCCTs). Our unsupervised approach uses a conditional generative adversarial network (GAN) for CTA to NCCT translation by generating a contrast map that allows for automatic extraction of vessel segmentations. Furthermore, we propose a 3D Frangi filter-based loss function to enhance tubular structures in the contrast map to improve vessel segmentations. We used a hold-out test set of 9 CTA volumes with manually annotated reference segmentations. We compared our unsupervised approach with a state-of-the-art supervised nnUnet, trained and evaluated with test set using 9-fold nested cross-validation. Evaluation metrics included voxel-wise Dice similarity coefficient (DSC), true positive rate (TPR), and false positive rate (FPR). The DSC was 4% lower for the unsupervised approach (DSC: 0.74) compared to the supervised nnUnet (DSC: 0.78). Both the TPR and FPR were higher for the unsupervised approach (TPR: 0.75, FPR/1000 voxels:2.05) compared to the supervised nnUnet (TPR:0.71, FPR/1000 voxels:0.87). Hence, the quantitative results showed that our unsupervised method approaches a supervised state-of-the-art segmentation network. The results demonstrate that an unsupervised generative deep learning approach for the segmentation of intracranial vessels is feasible without laborious manual segmentations. HighlightsO_LITo train supervised segmentation models laborious manual segmentations are needed C_LIO_LIUnsupervised generative deep learning does not require manual segmentations C_LIO_LIOur unsupervised method combines L1, adversarial, and a novel Frangiloss C_LIO_LIVarying loss function combinations can reduce false positives or false negatives C_LIO_LIOur method approached the performance of a state-of-the-art supervised nnUnet C_LI

14
Multimodal 3D Image Registration for Mapping Brain Disorders

Mahmood, H.; Islam, S. M. S.; Iqbal, A.

2024-08-26 neuroscience 10.1101/2024.08.24.609508 medRxiv
Top 0.1%
5.3%
Show abstract

We introduce an AI-driven approach for robust 3D brain image registration, addressing challenges posed by diverse hardware scanners and imaging sites. Our model trained using an SSIM-driven loss function, prioritizes structural coherence over voxel-wise intensity matching, making it uniquely robust to inter-scanner and intra-modality variations. This innovative end-to-end framework combines global alignment and non-rigid registration modules, specifically designed to handle structural, intensity, and domain variances in 3D brain imaging data. Our approach outperforms the baseline model in handling these shifts, achieving results that align closely with clinical ground-truth measurements. We demonstrate its efficacy on 3D brain data from healthy individuals and dementia patients, with particular success in quantifying brain atrophy, a key biomarker for Alzheimers disease and other brain disorders. By effectively managing variability in multisite, multi-scanner neuroimaging studies, our approach enhances the precision of atrophy measurements for clinical trials and longitudinal studies. This advancement promises to improve diagnostic and prognostic capabilities for neurodegenerative disorders.

15
Modular coupling of structure-function reveals network integration (rather than segregation) as the key mechanism for cognitive task discrimination

Fernandez Iriondo, I.; Jimenez Marin, A.; Aginako, N.; Zamora Lopez, G.; Erramuzpe, A.; Bonifazi, P.; Cortes, J.

2025-03-11 health informatics 10.1101/2025.03.10.25323674 medRxiv
Top 0.1%
5.2%
Show abstract

Understanding how structural and functional brain networks interact to support cognitive processes remains a central challenge in systems neuroscience. In this study, we investigate the dynamics of structure-function coupling (SFC) at the modular level across different cognitive tasks using multimodal neuroimaging data, including anatomical, diffusion, functional at rest and functional at different tasks. By constructing high-resolution structural and functional connectivity matrices, we assessed intra-modular (SFC-INT) and inter-modular (SFC-EXT) coupling to examine their roles in task-specific reorganization. Our results reveal that variations in SFC during cognitive tasks are primarily driven by changes in inter-modular coupling, emphasizing network integration over segregation. Specifically, tasks demanding higher cognitive flexibility, such as the gender stroop task, exhibited increased SFC-EXT, indicating enhanced integration between modules. In contrast, tasks focused on memory processing showed a tendency toward segregation, with lower SFC-EXT values. These findings highlight the significance of inter-modular integration as a flexible and dynamic mechanism underlying cognitive task discrimination. Our study advances the understanding of modular brain network dynamics, suggesting that the brains ability to integrate information across modules plays a pivotal role in cognitive flexibility and task performance.

16
Early Dementia Diagnosis in Older Adults through Machine Learning: A Cross-Sectional fMRI Data Analysis

Mostafa, F.; Sharma, K.; Khan, H.

2026-01-26 health informatics 10.64898/2026.01.24.26344772 medRxiv
Top 0.1%
5.0%
Show abstract

BackgroundEarly diagnosis of dementia can significantly improve care planning and patient outcomes while delaying progression. Machine learning algorithms can identify patterns in clinical and neuroimaging data that may aid in the early detection of dementia risk factors. ObjectiveTo evaluate the performance of the ensemble machine learning pipeline for classifying dementia status utilizing demographic, clinical, and imaging features, and to identify the most predictive variables contributing to model accuracy. MethodsA cross-sectional study analyzed 373 MRI scans from 150 subjects aged 60-98 years. Variables included cognitive scores (MMSE, CDR), volumetric brain measures (eTIV, nWBV, ASF), demographic features (age, sex, education), and socioeconomic status. After preprocessing and imputing missing values with random forests, tree-based variable selection was performed, and the dataset was split into training and test sets, with 5-fold cross-validation used for model validation. An ensemble of 8 machine learning models was used to classify patients as demented or non-demented. ResultsModel performance was assessed using the area under the receiver operating characteristic (ROC) curve (AUC), accuracy, sensitivity, specificity, precision, F1 score, and Matthews Correlation Coefficient (MCC). Random Forest achieved the highest AUC (0.963), while MLP demonstrated the highest accuracy (94.6%), F1-score (0.943), and MCC (0.893). CDR, MMSE, and ASF were identified as the top predictors. Performance was robust across folds in 5-fold CV, and feature importance analyses supported clinical relevance. ConclusionsEnsemble ML approaches offer high predictive performance in dementia classification. ML frameworks have the potential to be integrated into diagnostic support tools, enabling more accurate and earlier detection of dementia using clinical and imaging data.

17
A multimodal AI model for modeling the genetic risk factor of Alzeihmer's disease

Nguyen, T. M.; Woods, C.; Liu, J.; Wang, C.; Lin, A.-L.; Cheng, J.

2026-04-15 health informatics 10.64898/2026.04.13.26350803 medRxiv
Top 0.1%
4.8%
Show abstract

The apolipoprotein E{varepsilon} 4 (APOE4) allele is the strongest genetic risk factor for late-onset Alzheimers disease (AD), the most common form of dementia. APOE4 carriers exhibit cerebrovascular and metabolic dysfunction, structural brain alterations, and gut microbiome changes decades before the onset of clinical symptoms. Better understanding of the early manifestion of these physiological changes is critical for development of timely AD interventions and risk reduction protocols. Multi-modal datasets encompassing a wide range of APOE{varepsilon} 4 and AD associated biomarkers provide a valuable opportunity to gain insight into the APOE4 phenotype; however, these datasets often present analytical challenges due to small sample sizes and high heterogeneity. Here, we propose a two-stage multimodal AI model (APOEFormer) that integrates blood metabolites, brain vascular and structural MRI, microbiome profiles, and other clinical and demographic data to predict APOE4 allele status. In the first stage, modality-specific encoders are used to generate initial representa-tions of input data modalities, which are aligned in a shared latent space via self-supervised contrastive learning during pretraining. The contrastive learning objective encourages learning of informative and consistent representations across modalities through leveraging cross-modality relationships. In the second stage, the pretrained representations are used as inputs to a multimodal transformer that integrates information across modalities to predict a key AD-risk genetic variant (APOE4). Across 10 independent experimental runs with different train-validation-test splits, APOEFormer predicts whether an individual carries an APOE4 allele with an average prediction accuracy of 75%, demonstrating robust performance under limited sample sizes. Post hoc perturbation analysis of the predictive model revealed valuable insights into the driving components of the APOE4 phenotype-- including key blood biomarkers and brain regions strongly associated with APOE4.

18
Brain tumor MRI classification and identification using an image classification model via Convolutional Neural Networks

Mohanty, N.; Sarmadi, M.

2024-09-25 health informatics 10.1101/2024.09.13.23299832 medRxiv
Top 0.1%
4.7%
Show abstract

Malignant brain tumors are generally classified to be extremely aggressive and often can be fatal when not met with immediate action. Glioblastoma Multiforme is the most common type of malignant tumor found in the brain and is extremely aggressive. For this reason, advanced detection of malignant brain tumors is necessary for optimal mitigation. Conversely, the classification of tumors during Medical Resonance Imaging can be difficult due to bodily movements resulting in the movement of the tumor. The movement of the tumor can disrupt targeted radiotherapy and can also, at times, result in treatments about radiotherapy damaging healthy areas of the brain rather than areas of the tumor. This study proposes a novel deep learning system that can identify tumors from MRI images; which can be helpful for the case of early detection, as well as being able to track tumors during active imaging; resulting in higher efficiency with targeted radiotherapy. This is done utilizing Convolutional Neural Networks (CNNs) created via deep learning frameworks. With the image identification of tumors; 97% accuracy was achieved with optimization. The tumor-classification deep learning system achieved an accuracy of 98%. Further testing is required for optimization; with this optimization, higher accuracy can be reached.

19
PHOTONAI-Graph - A Python Toolbox for Graph Machine Learning

Ernsting, J.; Holstein, V. L.; Winter, N. R.; Sarink, K.; Leenings, R.; Gruber, M.; Repple, J.; Risse, B.; Dannlowski, U.; Hahn, T.

2023-06-29 health informatics 10.1101/2023.06.22.23291748 medRxiv
Top 0.1%
4.6%
Show abstract

Graph data is an omnipresent way to represent information in machine learning. Especially, in neuroscience research, data from Diffusion-Tensor Imaging (DTI) and functional Magnetic Resonance Imaging (fMRI) is commonly represented as graphs. Exploiting the graph structure of these modalities using graph-specific machine learning applications is currently hampered by the lack of easy-to-use software. PHOTONAI Graph aims to close the gap between domain experts of machine learning, graph experts and neuroscientists. Leveraging the rapid machine learning model development features of the Python machine learning API PHOTONAI, PHOTONAI Graph enables the design, optimization, and evaluation of reliable graph machine learning models for practitioners. As such, it provides easy access to custom graph machine learning pipelines including, hyperparameter optimization and algorithm evaluation ensuring reproducibility and valid performance estimates. Integrating established algorithms such as graph neural networks, graph embeddings and graph kernels, it allows researchers without significant coding experience to build and optimize complex graph machine learning models within a few lines of code. We showcase the versatility of this toolbox by building pipelines for both resting-state fMRI and DTI data in the hope that it will increase the adoption of graph-specific machine learning algorithms in neuroscience research.

20
DINMC: A Deep Learning Framework for Interpretable Normative Model Construction and Pathological Brain Alteration Detection

Ge, Z.; Liu, S.; Dou, W.

2026-05-29 bioinformatics 10.64898/2026.05.29.728652 medRxiv
Top 0.1%
4.5%
Show abstract

Background and ObjectiveNormative modeling is a key tool for understanding brain alterations in neurodegenerative diseases, such as cerebellar-type multiple system atrophy. However, existing methods lack interpretability and fail to capture clinically meaningful pathological changes. This study presents DINMC, a Deep Interpretable Normative Model Construction framework, which combines autoencoder-based learning with statistical hypothesis testing to better capture and interpret disease-specific neu-roanatomical changes. MethodsThe DINMC framework constructs normative models using neuroimaging data from multi-site large healthy cohorts. It utilizes a U-shaped convolutional autoencoder to train these models, which are then applied to reconstruct brain features from both patients and healthy controls within the same study cohort. Pathological confidence values are derived by fusing original and deviation feature spaces, offering a measure of disease-related pathology reflected in each dimension of the features. The framework was validated through statistical analysis and prognostic classification and regression tasks. ResultsThe pathological confidence provides valuable insights into the neuroanatomical regions most affected by the disease, as well as the correlation between changes in these regions and clinical assessment scales. Our optimal model outperform traditional methods in prognostic prediction tasks, with an AUC of 0.972 for classification tasks and an R2 of 0.432 for regression tasks. ConclusionDINMC provides a novel and interpretable framework for neuroimaging analysis. By combining deep learning and statistical hypothesis testing, this framework offers a unique solution to improving both the interpretability and performance of normative models in neuroimaging. The approach is scalable to other neuroimaging datasets, offering a versatile tool for broader biomedical applications.